Back

JCO Clinical Cancer Informatics

American Society of Clinical Oncology (ASCO)

Preprints posted in the last 7 days, ranked by how well they match JCO Clinical Cancer Informatics's content profile, based on 22 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Adaptive Post-Processing Recovers Most of the Gap to nnU-Net v2 in Head and Neck GTV Segmentation: A Paired Three-Arm HECKTOR 2025 Benchmark

Oyarzun Silva, R.; Hernandez Hernandez, P.

2026-08-31 radiology and imaging 10.64898/2026.08.28.26361649 medRxiv
Top 0.1%
11.7%
Show abstract

Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.

2
Pretrained transformers applied to population cancer registries improve survival prediction in label-scarce and previously unseen cancers

Gao, Y.; Yu, S.; Xia, Y.; Chen, S.; Xia, S.; An, R.; Zeng, J.; Zhao, F.; Ma, Y.; Wang, Y.; Xie, X.; Zhang, J.

2026-09-03 oncology 10.64898/2026.08.30.26361693 medRxiv
Top 0.1%
6.6%
Show abstract

Prognostic models in oncology are developed one cancer at a time, from that cancer's own labelled outcomes, and fail where prognostic information is scarcest. Rare cancers account for roughly a fifth of diagnoses and most paediatric malignancies, yet seldom supply enough events for a reliable time-to-event model. We therefore asked whether a representation learned without outcome labels can supply what those cohorts cannot. A Transformer encoder was pretrained by masked field-value modelling on 9425135 tumour records from the SEER 17 registries, diagnosed in 2000 to 2023. Only diagnosis-time fields passing a fail-closed coding-verification gate were admitted, and each record was emitted as an era-specific and a harmonised view, keeping two decades of recoding auditable. The encoder was then frozen and read by a linear Cox head for overall survival. Nine rare cancers were removed from the pretraining corpus entirely, each requiring an independent pretraining run. On a sealed test partition, all nine exceeded an architecture-identical random frozen encoder in Harrell concordance by +0.0034 to +0.0368, every lower confidence limit above zero. At 256 labelled patients, all 67 cancers favoured the pretrained representation over budget-matched Cox regression, median difference +0.0283. The advantage was bounded: given the entire training set, Cox regression was favoured in seven of nine rare cancers. The encoder did not outperform a field-frequency baseline on its own objective, so upstream reconstruction did not predict downstream transfer. Outcome-agnostic registry pretraining carries prognostic signal into cancers it has never seen, and is most useful where labels are fewest, without establishing clinical utility.

3
PCGS: biomarker and risk group identification for Pediatric Cancers via explainable Graph neural networks with Shapley values

Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.

2026-09-01 health informatics 10.64898/2026.08.27.26361540 medRxiv
Top 0.1%
6.1%
Show abstract

Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.

4
Genome Profiling of Actionable Cancer Targets (NYU LG-PACT) for Clinical Patient Molecular Diagnostics and Treatment

Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.

2026-09-01 oncology 10.64898/2026.08.27.26361341 medRxiv
Top 0.2%
3.3%
Show abstract

Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.

5
Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome

Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.

2026-08-31 oncology 10.64898/2026.08.26.26361432 medRxiv
Top 0.2%
3.2%
Show abstract

Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.

6
RedFuMOS: A novel approach for multi-omics and clinical data-driven patient stratification

De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.

2026-08-31 health informatics 10.64898/2026.08.26.26361415 medRxiv
Top 0.3%
2.8%
Show abstract

Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.

7
AURORA: Analysing and understanding responses to oncological regimens with artificial intelligence

Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.

2026-09-02 health informatics 10.64898/2026.08.30.26361778 medRxiv
Top 0.4%
1.7%
Show abstract

Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.

8
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 0.4%
1.5%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

9
Spatial mapping of loco regional recurrences and disease related outcomes in breast cancer patients with Internal Mammary Node (IMN) positivity at presentation treated with a curative intent using moderately hypo-fractionated radiotherapy

Chowdhury, D.; Chatterjee, S.; Chakraborty, S.; Mahata, A.; Vashistha, B.

2026-09-03 oncology 10.64898/2026.08.31.26361095 medRxiv
Top 0.5%
1.3%
Show abstract

Purpose/Objective There is paucity of data reporting outcomes of breast cancers with initial internal mammary nodal involvement and no visceral metastases, treated with curative hypofractionated radiotherapy . We report the outcomes from a tertiary centre alongside spatial patterns of recurrences in the above group Material/Methods For this retrospective cross-sectional study, consecutive patients contoured as per the ESTRO 2013 guidelines, treated between 2016-2022 were eligible if their diagnostic imaging demonstrated involvement of the internal mammary nodes. Radiotherapy (40 Gy/15#/3 weeks) was delivered to the residual breast / thoracic wall, SCF region corresponding to the ESTRO lymph node level 4 and internal mammary chain nodes. Residual IMN/ level 4 nodes received a boost of 10Gy/5#. Spatial mapping of sites of recurrence at the local site and three nodal sites (axilla, SCF and IMN) was performed using deformable image registration. Sites of recurrence at the local site and three nodal levels were contoured separately. Volumetric intersection of the recurrent gross tumour volume (GTV_recurrence) with treated clinical target volume (CTV) was calculated. Actuarial overall (OS), disease free survival (DFS) & cumulative incidence of local (LR), regional (RR) and loco-regional recurrence(LRR) were calculated using Kaplan Meier method. Univariate comparison of outcomes with or without residual disease was performed using the log rank test. Results The median age of the 61 eligible women was 49 years. 77% received neoadjuvant chemotherapy and the rest adjuvant chemotherapy. 82% patients had a mastectomy. Axillary lymph node dissection was done in 96.7%. Boosts to residual IMN and SCF nodes were delivered to 21(34.4%) and 2 (3.3%) respectively. Median follow up was 3.6 years. Out of the 61 patients, 42 patients were disease free with an estimated 3 year disease free survival of 75% (95% CI 64, 88%). Spatial mapping of locoregional recurrence was possible in all but 1 patient with local (only) recurrence who was lost to follow-up after mammogram only. Among the patients with loco regional recurrence 1 had recurrence in local site + SCF +axilla, 3 had recurrence in the SCF+axilla, 2 in the SCF+IMN and 1 in the axilla+SCF+IMN. Only one patient had isolated axillary recurrence or isolated SCF recurrence. There were no IMN only recurrences. Among the 8 patients with nodal recurrence, a total of 27 individual GTV_recurrence were identified in the axilla(n=11), SCF(n=11) and IMN (n=5). IMN recurrences showed complete or partial overlap with CTV. SCF recurrences were a mix with predominantly in-field recurrences while axillary recurrences occurred outside the treated volume.Four (6.6%) patients had Grade 2 lymphoedema as documented late side effect. Conclusion Aggressive treatment of IMN disease with adjuvant radiation is effective with good locoregional control. Systemic recurrences are common and may benefit from intensification strategies.

10
Augmenting Deep Learning-Based PSMA PET/CT Metastasis Segmentation with a Population-Level Spatial Atlas

Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26361439 medRxiv
Top 0.6%
1.1%
Show abstract

Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.

11
Surprisal-based large language models reveal immunologic insights in lobular breast cancer

Majumder, B. P.; Linak, J. A.; Adamson, R.; Aguilera, R. L.; Agarwal, D.; Reitz, Z.; Loiselle, S.; Devarakonda, S.; Clark, P.; Paulson, K. G.; Stanton, S.

2026-08-31 oncology 10.64898/2026.08.25.26361365 medRxiv
Top 0.6%
1.0%
Show abstract

In large data sets discovery is often limited to pre-conceived hypotheses and data fishing. Here we tested whether systematic exploration of AI generated hypotheses could uncover clinically meaningful signals in extensively studied data. We deployed AutoDiscovery, a newly launched large language model (LLM) framework designed to search for hypotheses based on surprisal and systematically interrogate complex datasets, on The Cancer Genome Atlas breast cancer cohort. The system did not identify clinically meaningful novel findings without human input. However, a seeded warm-start run with minimal text input from an oncologist revealed multiple interesting and surprising hypotheses. Among these was that a robust immune signature was present across all subtypes of invasive lobular carcinoma (ILC) that exceeded invasive ductal carcinoma (IDC). This observation was independently validated in independent cohorts and confirmed by high-sensitivity multi-immunofluorescence tumor tissue analyses. These results suggest immunotherapy approaches should be tested in ILC including early stage ER+HER2- ILC; these patients are currently excluded from large neoadjuvant immunotherapy trials. They further demonstrate that surprisal-based hypothesis generation frameworks can extract previously unappreciated patterns from deeply interrogated cancer datasets and imply that disease domain experts working with LLMs can derive more meaningful insights from complex data than either could achieve alone.

12
Psychosocial Stress and Allostatic Load Among Underrepresented Minority Women with Familial Cancer Risk

Shachar, E. K.; Haas, R.; Rodriguez, V. E.; Lester, J.; Siavoshi, M. A.; Kwan, L.; Niell-Swiller, M.; Spellman, P. T.; Boutros, P. C.; Chang, V. Y.; Karlan, B. Y.

2026-08-31 public and global health 10.64898/2026.08.26.26361226 medRxiv
Top 0.7%
0.9%
Show abstract

Importance: Chronic stress may contribute to adverse health outcomes through cumulative physiologic dysregulation. Allostatic load (AL), a composite measure of multisystem physiologic burden, may capture biologic effects of structural, social, and psychosocial stress not reflected by self-reported measures. Objective: To evaluate racial and ethnic differences in AL among women with familial cancer risk and examine how socioeconomic status, psychosocial factors, clinical characteristics, and health behaviors contribute to variations in AL. Design: Cross-sectional study of underrepresented minority participants enrolled in the HERSTORY cohort from October 2023 through September 2025, with comparison participants from the UCLA ATLAS biobank. Setting: UCLA academic health system. Participants: The study included 303 racially and ethnically diverse female HERSTORY participants aged [&ge;]35 years with a family history of cancer and matched non-Hispanic White female ATLAS participants (n=709). Exposures: Race and ethnicity, age, neighborhood deprivation, cancer history and stage, depression, perceived stress, cancer worry, and physical activity. Main Outcomes and Measures: The primary outcome was AL, calculated from cardiometabolic and organ-function measures. A secondary index incorporated race- and ethnicity-specific neutrophil-to-lymphocyte ratio (NLR) derived from 326,826 women in the UCLA Health population. Multivariable regression models evaluated factors associated with elevated AL. Results: Compared with matched non-Hispanic White participants, Black and Asian/Pacific Islander HERSTORY participants had significantly higher AL after adjustment. Hispanic/Latina participants did not have significantly elevated AL. Older age, greater area-level socioeconomic deprivation, and depression were independently associated with higher AL. Prior cancer diagnosis, cancer worry and perceived stress were not significantly associated with AL, whereas regular physical activity was associated with lower AL. Among cancer patients, advanced stage was associated with greater AL. Conclusions and Relevance: This study demonstrates elevated AL among understudied racial/ethnic minority groups with familial cancer risk and identifies associations with neighborhood deprivation, depression, and physical activity. The association between cancer stage and AL suggests that physiologic stress may reflect variation in cancer burden. The lack of association with perceived stress and cancer worry further indicates that physiologic and self-reported psychosocial measures capture distinct dimensions of stress. The development of race/ethnicity-specific NLR thresholds derived from large population samples provide a benchmark for future studies.

13
Global Adoption of openEHR Clinical Data Repositories: A Vendor and Community Survey

Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.

2026-08-31 health informatics 10.64898/2026.08.27.26361529 medRxiv
Top 0.8%
0.8%
Show abstract

The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.

14
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.8%
0.6%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

15
LLM-assisted evidence audit of late-stage cancer incidence as a screening trial endpoint

Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.

2026-08-31 oncology 10.64898/2026.08.29.26361733 medRxiv
Top 0.9%
0.6%
Show abstract

Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.

16
Soft-Tissue versus Hematologic Primary Malignant Cardiac Tumors: Demographics and First-Course Treatment Patterns in the SEER Registry

Mathew, Z.; Mehta, R.; Kim, S.; Jeyaraj, J.; Asif, T.

2026-08-31 cardiovascular medicine 10.64898/2026.08.25.26361262 medRxiv
Top 0.9%
0.5%
Show abstract

Background: Primary malignant cardiac tumors (PMCTs) are rare and histologically heterogeneous. Objective: To compare demographics, specific ICD-O-3 morphologies, first-course treatment patterns, annual registered case counts, and unadjusted overall survival between soft-tissue and hematologic PMCTs. Methods: We identified 730 PMCT cases diagnosed from 2000 to 2021 in SEER 18 (ICD-O-3 topography C38.0). Histologic lineage was assigned from ICD-O-3 morphology. Comparative analyses included soft-tissue (n=458) and hematologic (n=212) tumors. First-course variables were primary-site surgery, chemotherapy (yes versus no/unknown), and radiotherapy (radiation versus none/unknown). Groups were compared with chi-square tests. Overall survival was estimated with Kaplan-Meier methods; follow-up was truncated at 120 months. Results: Soft-tissue PMCTs occurred predominantly at ages 45-64 years (67.9%), whereas hematologic PMCTs occurred predominantly at age [&ge;]65 years (63.2%; p<0.001). Men comprised 59.9% of hematologic and 49.3% of soft-tissue cases (p=0.014). The leading soft-tissue morphology was hemangiosarcoma/angiosarcoma (ICD-O-3 9120/3; 201/458, 43.9%); synovial sarcoma accounted for 20/458 cases (4.4%). Diffuse large B-cell lymphoma, NOS, accounted for 131/212 hematologic tumors (61.8%). Any primary-site surgery was recorded in 66.6% of soft-tissue versus 15.6% of hematologic cases (p<0.001). Chemotherapy was recorded in 67.5% versus 51.1% (p<0.001), and radiotherapy in 9.0% versus 20.5% (p<0.001). In exploratory Kaplan-Meier analyses, hematologic patients with recorded chemotherapy had higher unadjusted 120-month overall survival than those without recorded chemotherapy (42.0% versus 12.2%; log-rank p=7.5x10-). Radiation-associated survival differences were not statistically significant in either lineage. Conclusions: Soft-tissue and hematologic PMCTs have distinct age distributions, named histologies, and first-course treatment patterns in SEER. These findings describe registry coding and do not establish treatment effectiveness or population incidence.

17
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 1%
0.5%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

18
Cost-Aware Active Feature Acquisition for Differential Diagnosis under Realistic Clinical Availability Constraints

Bingham, J. C.; Arussy, N.

2026-08-31 health informatics 10.64898/2026.08.30.26361745 medRxiv
Top 1%
0.5%
Show abstract

Active Feature Acquisition (AFA) adaptively selects which diagnostic test to order next and offers a route to reduce unnecessary laboratory testing in acute care. Existing clinical AFA evaluations, however, assume every feature can be retrieved on demand and split data at the visit level, both of which inflate apparent performance. We re-evaluate cost-aware AFA under constraints designed to reflect deployment. From MIMIC-IV we constructed a cohort of 64,766 acute admissions (39,884 patients; 21 conditions; 55 features in 30 test panels) with a patient-level split, a 12-hour decision cutoff, and a per-patient availability mask from what was actually measured, and priced panels using the 2026 Medicare fee schedule under panel-level billing. We evaluated EIG-Cost, which scores each panel by Monte-Carlo Expected Information Gain penalised by its dollar cost, against eight published methods across budgets \30--$60 over five patient-level resamples. At a $30 budget, EIG-Cost achieved the highest macro-F1 (0.188, 95% CI [0.185, 0.191]) at the lowest cost ($17.28), exceeding the strongest baseline in all five resamples (p<0.001; Cohen's d=4.0), and led at every budget. Three of the eight methods collapsed to a vitals-only baseline (macro-F1 approx 0.040), acquiring nothing even at higher budgets, a genuine failure to adapt to availability rather than a budget limitation. Despite modest absolute accuracy, EIG-Cost's probabilities were well-calibrated (expected calibration error $0.048$). Under realistic availability constraints, clinical AFA is substantially harder than full-availability benchmarks imply, several published methods fail outright, and cost-aware information-gain scoring is a robust choice in this harder setting.

19
Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology

Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.

2026-09-02 hematology 10.64898/2026.09.01.26361881 medRxiv
Top 1%
0.4%
Show abstract

Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.

20
CPT/HCPCS Code Recommendation from Clinical Notes: A Comparative Evaluation of AI Methods

Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.

2026-08-31 health informatics 10.64898/2026.08.29.26361731 medRxiv
Top 1%
0.4%
Show abstract

Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.